Papers by Aditya Sanjiv Kanade
Mind’s Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluations of multimodal large language models (MLLMs) have demonstrated compelling visual understanding in recent years. |
| Approach: | They propose a multimodal large language model with eight visuo-cognitive tasks inspired by classic human intelligence tests organized under a novel A–R–T taxonomy: Abstraction, Relation, and Transformation. |
| Outcome: | The proposed frameworks are based on eight visuo-cognitive tasks inspired by human intelligence tests and organized under a novel A–R–T taxonomy: Abstraction, Relation, and Transformation. |
MIR: Methodology Inspiration Retrieval for Scientific Research Problems (2025.acl-long)
Copied to clipboard
Aniketh Garikaparthi, Manasi Patwardhan, Aditya Sanjiv Kanade, Aman Hassan, Lovekesh Vig, Arman Cohan
| Challenge: | Existing methods for generating ideas rely on grounding the discovery process within the literature, but their effectiveness varies significantly with the quality and nature of the retrieved literature. |
| Approach: | They construct a methodological inspiration retrieval task using a citation-based methodology adjacency graph and embed an "intuitive prior'' into dense retrievers. |
| Outcome: | The proposed method achieves significant gains in Recall@3 and mAP over strong baselines. |
Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs (2026.eacl-long)
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) show reasoning promise, yet their visual perception is a bottleneck. |
| Approach: | They propose a visual perception benchmark to test the visual perception of MLLMs. |
| Outcome: | The proposed benchmark examines MLLMs' visual perception abilities with 1758 images and 2612 questions. |
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs (2026.acl-short)
Copied to clipboard
| Challenge: | Existing multimodal reasoning models lack generalized spatial intelligence, a new study shows . a critical gap exists in the field of vision-centric reasoning, the authors argue . |
| Approach: | They evaluate 16 multimodal reasoning models using Chain-of-Though (CoT) based thinking . they find that CoT prompting consistently degrades performance in visual spatial reasoning . |
| Outcome: | The proposed model hallucinates visual details from textual priors even when the image is absent. |